Papers with Rhetorical Structure Theory
GCDT: A Chinese RST Treebank for Multigenre and Multilingual Discourse Parsing (2022.aacl-short)
Copied to clipboard
| Challenge: | GCDT is the largest hierarchical discourse treebank for Mandarin Chinese in the framework of Rhetorical Structure Theory (RST). |
| Approach: | They propose to use a Chinese hierarchical discourse treebank to parse Mandarin Chinese using relation inventory and a multilingual training program. |
| Outcome: | The proposed dataset includes state-of-the-art scores for Chinese RST parsing and RST Parsing on the English GUM dataset, using cross-lingual training in Chinese and English with multilingual embeddings. |
Extractive Summarisation for German-language Data: A Text-level Approach with Discourse Features (2022.coling-1)
Copied to clipboard
| Challenge: | Using RST, extractive summarisation involves using select phrases and sentences as a summary, which still remains a strong method for producing summaries despite its simple nature. |
| Approach: | They propose to use RST-based features to analyse the connection between summary sentences and several RST features and transfer these insights to various automated summarisation models. |
| Outcome: | The proposed models are based on the best features proposed over the last 20+ years and incorporate the best ones into the proposed models. |
Neural RST-based Evaluation of Discourse Coherence (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing discourse parsers cannot predict coherent texts without using silver-standard features. |
| Approach: | They propose a tree-recursive neural model which takes advantage of the text’s RST features produced by a state of the art RST parser and compares it to the current state of art. |
| Outcome: | The proposed model achieves state-of-the-art accuracy on the Grammarly Corpus for Discourse Coherence (GCDC) and has 62% fewer parameters than existing models. |
Improving Neural RST Parsing Model with Silver Agreement Subtrees (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods for Rhetorical Structure Theory (RST) parsing use supervised learning, but the RST-DT is small due to the costly annotation of RST trees. |
| Approach: | They propose to use silver data to improve RST parsing models by using annotated silver data. |
| Outcome: | The proposed method achieves the best micro-F1 scores for Nuclearity and Relation at 75.0 and 63.2 . it also achieves a remarkable gain in relation score against the previous state-of-the-art parser. |
RST Parsing from Scratch (2021.naacl-main)
Copied to clipboard
| Challenge: | Fig. 1 shows a document level discourse parser that performs top-down end-to-end parsing without requiring segmentation . |
| Approach: | They propose a top-down end-to-end formulation of document level discourse parsing in the Rhetorical Structure Theory framework. |
| Outcome: | The proposed model outperforms existing methods in end-to-end parsing and parse with gold segmentation without handcrafted features. |
RSTGen: Imbuing Fine-Grained Interpretable Control into Long-FormText Generators (2022.naacl-main)
Copied to clipboard
| Challenge: | Using a framework based on Rhetorical Structure Theory, we aim to improve the cohesion and coherence of long-form text generated by language models. |
| Approach: | They propose a framework that utilises Rhetorical Structure Theory to control the discourse structure, semantics and topics of generated text. |
| Outcome: | The proposed framework performs competitively against existing models while offering significantly more controls over generated text than alternative methods. |
Can we obtain significant success in RST discourse parsing by using Large Language Models? (2024.eacl-long)
Copied to clipboard
| Challenge: | Experimental results show that LLMs with tens of billion parameters can perform discourse parsing tasks. |
| Approach: | They employ Llama 2 and fine-tune it with QLoRA to achieve similar results . they show that LLMs with tens of billion parameters can perform a wide range of NLP tasks . |
| Outcome: | The proposed model performs better than existing models on three benchmark datasets. |
End-to-End Argument Mining over Varying Rhetorical Structures (2023.findings-acl)
Copied to clipboard
| Challenge: | Rhetorical Structure Theory implies no single discourse interpretation of a text . inconsistent parsing of similar structures can result in inconsistent argumentation analysis . |
| Approach: | They propose a deep dependency parsing model to assess the connection between rhetorical and argument structures. |
| Outcome: | The proposed model allows for end-to-end argumentation analysis using a rhetorical tree instead of a word sequence. |
Beyond the Final Actor: Modeling the Dual Roles of Creator and Editor for Fine-Grained LLM-Generated Text Detection (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods to detect large language models (LLMs) use binary or ternary classifications, which can only distinguish pure human/LLM text or collaborative text at best. |
| Approach: | They propose a fine-grained method that characterizes distinct signatures of creator and editor by using Rhetorical Structure Theory to construct a logic graph for creator's foundation and extracting Elementary Discourse Unit (EDU)-level features for the editor's style. |
| Outcome: | The proposed method outperforms 12 baselines in identifying fine-grained types with low false alarms, offering a policy-aligned solution for LLM regulation. |
A Multi-layer Annotated Corpus of Argumentative Text: From Argument Schemes to Discourse Relations (L18-1)
Copied to clipboard
| Challenge: | Recent interest in Argumentation Mining has brought to the fore the need for corpora annotated with argument information, which can be used as training data. |
| Approach: | They propose a set of guidelines for the annotation of argument schemes and a new annotation tool for the 'inferential' argument schemes. |
| Outcome: | The proposed corpus includes 112 argumentative microtexts and a new annotation tool. |
Developing the Bangla RST Discourse Treebank (L18-1)
Copied to clipboard
| Challenge: | a corpus in Bangla is annotated for coherence relations between text segments representing propositions . the corpus is a valuable resource for conducting discourse studies for Bangla . |
| Approach: | They propose to build a Bangla-annotated corpus which includes 266 Bangla texts . they use Rhetorical Structure Theory as the theoretical framework to develop the corpus . |
| Outcome: | The proposed corpus contains 266 Bangla texts annotated for coherence relations . the research could be used for discourse studies and for developing NLP applications . |
AMPERSAND: Argument Mining for PERSuAsive oNline Discussions (D19-1)
Copied to clipboard
| Challenge: | Argument mining is a field of corpus-based discourse analysis that involves the automatic identification of argumentative structures in text. |
| Approach: | They propose a computational model for argument mining in online persuasive discussion forums that brings together the micro-level (argument as product) and macro-level models of argumentation. |
| Outcome: | The proposed model improves on existing models using pointer networks and a pre-trained language model. |
MCDTB: A Macro-level Chinese Discourse TreeBank (C18-1)
Copied to clipboard
| Challenge: | Discourse analysis is becoming increasingly important in the field of natural language processing. |
| Approach: | They propose to annotate macro discourse information and additional discourse information to make annotation more objective and accurate. |
| Outcome: | The results show that the annotations are more objective and accurate than the previous ones. |
Enhancing the AI2 Diagrams Dataset Using Rhetorical Structure Theory (L18-1)
Copied to clipboard
| Challenge: | Existing annotation schemas for diagrams are based on Rhetorical Structure Theory (RST) paper documents proposed schema, reports on inter-annotator agreement for this task, and discusses use of AI2D-RST for research on multimodality and artificial intelligence. |
| Approach: | They propose to replace the annotation of semantic relations between diagram elements by building on Rhetorical Structure Theory (RST) the paper documents the proposed annotation schema, describes challenges in applying RST to diagrams, and reports on inter-annotator agreement for this task. |
| Outcome: | The proposed schema is based on Rhetorical Structure Theory, which has been used to describe the multimodal structure of diagrams and documents. |
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization (2023.acl-long)
Copied to clipboard
| Challenge: | Contemporary leading-edge systems for abstractive (long) text summarization employ Transformer encoderdecoder architectures that only consider the nuclearity annotation . |
| Approach: | They propose to incorporate Rhetorical Structure Theory into a novel summarization model that incorporates both the types and uncertainty of rhetorical relations. |
| Outcome: | The proposed model outperforms state-of-the-art models on automatic metrics and human evaluation. |
TreeAnnotator: Versatile Visual Annotation of Hierarchical Text Relations (L18-1)
Copied to clipboard
| Challenge: | TREEANNOTATOR is a browser-based tool for annotating tree-like structures . it provides a wider range of formats and provides graphical annotations . |
| Approach: | They evaluate TREEANNOTATOR, a browser-based tool for annotating tree-like structures, in particular structures that jointly map dependency relations and inclusion hierarchies, as used by Rhetorical Structure Theory. |
| Outcome: | The GUI interface is user-friendly and provides two visualization modes. |
Using Discourse Information for Education with a Spanish-Chinese Parallel Corpus (L18-1)
Copied to clipboard
| Challenge: | Discourse information is crucial for many NLP tasks due to the great distance that spans between the two languages. |
| Approach: | They propose to use a Spanish-Chinese parallel corpus with annotated discourse information to serve for bilingual language education. |
| Outcome: | The proposed corpus is composed of 100 Spanish-Chinese parallel texts, and all the discourse markers (DM) have been annotated to form the education source. |
A Unified Linear-Time Framework for Sentence-Level Discourse Parsing (P19-1)
Copied to clipboard
| Challenge: | a new neural framework for sentence-level discourse analysis is proposed . a discourse segmenter and a parser are based on pointer networks and operate in linear time . |
| Approach: | They propose a neural framework for sentence-level discourse analysis in accordance with Rhetorical Structure Theory . they use a discourse segmenter and a parser to construct a discursive tree in a top-down fashion . |
| Outcome: | The proposed framework surpasses previous approaches on both tasks and human agreement on both. |
Developing a Rhetorical Structure Theory Treebank for Czech (2024.lrec-main)
Copied to clipboard
| Challenge: | a paper on the Czech RST Discourse Treebank is the first version of a textual annotation system based on the Rhetorical Structure Theory . document is annotated using the RST, a global coherence model proposed by Mann and Thompson . |
| Approach: | They introduce the first version of the Czech RST Discourse Treebank . paper presents an annotation process and provides corpus statistics and evaluation . |
| Outcome: | The paper presents the first version of the Czech RST Discourse Treebank . the treebank includes two gold annotations representing divergent interpretations . |
Bilingual Rhetorical Structure Parsing with Large Parallel Annotations (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing large RST corpora are inconsistent in annotation guidelines, genre representation, source selection, and relation definitions. |
| Approach: | They propose a parallel Russian annotation for a large and diverse English GUM RST corpus. |
| Outcome: | The proposed RST parser achieves state-of-the-art results on English and Russian corpus . it demonstrates effectiveness in monolingual and bilingual settings, transferring even with limited second-language annotation. |
Video Discourse Parsing and Its Application to Multimodal Summarization: A Dataset and Baseline Approaches (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Fig. 1 shows the video's story structure and event relationships in discourse parsing. |
| Approach: | They propose to construct an RST tree for a video to represent its storyline and illustrate the event relationships between events. |
| Outcome: | The proposed model outperforms two existing approaches to video RST parsing: the ‘parsing after captioning’ framework and parser using visual features. |
Split or Merge: Which is Better for Unsupervised RST Parsing? (D19-1)
Copied to clipboard
| Challenge: | Rhetorical Structure Theory (RST) parsers have been based on supervised learning approaches that require an annotated corpus of sufficient size and quality. |
| Approach: | They propose two unsupervised methods that build an optimal RST tree based on a dissimilarity score function for splitting a text span into smaller ones and a similarity score for merging two adjacent spans into a large one. |
| Outcome: | The proposed method achieves the best score on English and German RST treebanks, around 0.8 F1 score, close to the previous supervised parsers. |
Multilingual Neural RST Discourse Parsing (2020.coling-main)
Copied to clipboard
| Challenge: | Existing studies on text discourse parsing for English are limited due to the lack of annotated data. |
| Approach: | They propose to use multilingual vector representations and segment-level translation to establish a neural, cross-lingual discourse parser. |
| Outcome: | The proposed model achieves state-of-the-art on cross-lingual, document-level discourse parsing on all sub-tasks. |
AMALGUM – A Free, Balanced, Multilayer English Web Corpus (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of 4M tokens is available online with a large number of high-quality annotation layers. |
| Approach: | They propose to use a genre-balanced English web corpus with multiple annotation layers . they harness knowledge from multiple annotation layer to achieve a "better than NLP" benchmark . |
| Outcome: | The proposed corpus is genre-balanced and features high-quality automatic annotation layers. |
Coherence of Argumentative Dialogue Snippets: A New Method for Large Scale Evaluation with an Application to Inference Anchoring Theory (2025.findings-emnlp)
Copied to clipboard
| Challenge: | illocutionary acts and propositional relations impact dialogue coherence, whereas propositional acts do not. |
| Approach: | They propose a method for testing the components of theories of dialogue coherence through utterance substitution and apply it to Inference Anchoring Theory (IAT) |
| Outcome: | The proposed method is applied to 933 dialogue snippets and 87 annotators. |
Discourse Realization of Generics in Human and LLM-generated Texts (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models produce texts that appear coherent and credible, even when their factual reliability is uncertain. |
| Approach: | They propose a text-level genericity score derived from clause-level annotations and apply it to argumentative essays produced by humans and LLMs. |
| Outcome: | The proposed model is less generic than LLM-produced arguments, the study shows . higher genericity correlates with less structured, paratactic structures, the research shows a. |
Cross-Document Cross-Lingual NLI via RST-Enhanced Graph Fusion and Interpretability Prediction (2025.emnlp-main)
Copied to clipboard
| Challenge: | Despite the development of many subdirections, Cross-Document Cross-Lingual NLI remains largely unexplored. |
| Approach: | They propose a novel paradigm that extends traditional NLI capabilities to multi-document, multilingual scenarios by integrating RST-enhanced graph fusion with interpretability-aware prediction. |
| Outcome: | The proposed method improves on existing models and document-level NLI to multi-document, multilingual scenarios. |